Back

Molecular Ecology Resources

Wiley

Preprints posted in the last 7 days, ranked by how well they match Molecular Ecology Resources's content profile, based on 171 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.

1
AmPair: automating housekeeping-gene primer design for species-level metataxonomics

Xu, X.; Yang, X.

2026-09-01 bioinformatics 10.64898/2026.08.25.746527 medRxiv
Top 0.3%
7.7%
Show abstract

Amplicon sequencing of the 16S rRNA gene is the most widely used approach for profiling bacterial communities, but its taxonomic resolution is typically limited to the genus level. Many species carry multiple divergent 16S rRNA alleles that overlap across species boundaries, an ambiguity that even full-length, long-read sequencing cannot fully resolve. Shotgun metagenomics achieves species-level resolution but remains costly, particularly when only a single genus is of interest. Amplicon sequencing of rapidly evolving, protein-coding housekeeping genes offers a cost-effective alternative, yet no tool exists to identify suitable primer sets for a given target taxon. Here we present AmPair, a Snakemake pipeline that, given a target genus and one or more candidate housekeeping genes, designs and ranks primer pairs binding conserved regions while flanking a variable region capable of species-level discrimination, and validates them in silico across all available genomes. Using the genus Bacillus and the housekeeping gene tuf as a case study, the primer set recommended by AmPair amplified 99% of 2,392 genomes; only 0.04% carried multiple alleles and none showed inter-species allele overlap, compared with 91.41% and 69.49%, respectively, for the standard 16S rRNA V1-V9 region. Applied to a Bacillus community profiled by Nanopore sequencing, the same primers resolved closely related species. AmPair thus offers a generalizable and accessible route to species-level community profiling.

2
Rapid isothermal amplification of diatom rbcL from eDNA and eRNA reveals their abundance and photosynthetic physiology

Verret, F. G.; Hartle-Mougiou, K.; Chantzaras, C.; Peltekis, A.; Margiotta, F.; Sarno, D.; Cardini, U.; Alba, M.; Pizziol, V.; Markopoulos, I.; Papadopoulou, I.; Percopo, I.; Tramontano, F.; Maselli, M.; Novellino, A.; Psarra, S.; Montresor, M.; Mowlem, M. C.; Gizeli, E.; Valiadi, M.

2026-08-31 microbiology 10.64898/2026.08.30.748096 medRxiv
Top 0.3%
6.8%
Show abstract

Diatoms are major contributors to marine primary production, yet current approaches for monitoring their abundance and function rely on coarse satellite chlorophyll estimates or sparse cell count and carbon fixation measurements. Molecular markers are a promising approach for high-resolution measurement of both abundance and metabolic activity through analysis of environmental DNA (eDNA) and RNA (eRNA). We present an isothermal quantitative recombinase polymerase amplification (qRPA) assay targeting rbcL gene copies and transcripts of marine diatoms, operating at low temperature and producing results in less than 15 min. We demonstrate specificity and calibration across diverse diatom taxa, then apply the assay to eDNA and eRNA samples from the Mare Chiara Long-Term Ecological Research site in the Bay of Naples, Italy, alongside microscopy, chlorophyll, physicochemical, and carbon-fixation data. Diatom rbcL DNA tracked abundance across five orders of magnitude despite seasonal shifts in community composition. Combining molecular and optical data revealed increased cellular rbcL copies and chlorophyll in low-light winter populations, suggesting enhanced photosynthetic capacity despite lower abundance. Furthermore, rbcL RNA reflected total carbon fixation rates and identified populations with differing carbon fixation activity. These results support rapid, RPA-based rbcL quantification as a robust approach for biomolecular ocean observing.

3
High-Molecular-Weight Genomic DNA Extraction from Recalcitrant Australian Plants: An Optimised CTAB Protocol for Anigozanthos

Rajput, R.; Saha, L.; Ahmed, Z.; Naiker, P.; Do, L.; Bisset, A.; Hooper, C.

2026-08-31 plant biology 10.64898/2026.08.29.741951 medRxiv
Top 0.4%
5.6%
Show abstract

High-phenolic plant genera present a major technical limitation in genomic research. Standard extraction approaches that perform reliably across diverse flora often perform poorly when applied to recalcitrant taxa, producing low DNA yield and integrity incompatible with sequencing requirements. The genus Anigozanthos (Kangaroo paws) from the family Haemodoraceae exemplifies this problem. We identified key physicochemical factors governing extraction failure in this genus and resolved them through targeted modifications to lysis chemistry and contaminant management. The resulting protocol achieved a near threefold improvement in DNA purity, substantially reducing contaminant carry over and consistently yielded high-integrity, long DNA fragments (DIN > 7) across a diverse sample set spanning cultivated and wild material across four diverse genera of Haemodoraceae. We also tested a straightforward purity assessment framework that can be implemented in any standard molecular laboratory, enabling rapid pre-submission quality assessment without the need for specialised equipment. Together these advances open a practical path to genomic characterisation of Anigozanthos that establishes a transferable model for genomic research across Australia ' s chemically complex native flora.

4
Rclade: automated taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R

Zeng, Z.; Wang, Y.

2026-09-01 bioinformatics 10.64898/2026.08.27.747462 medRxiv
Top 1%
2.1%
Show abstract

Background: Reproducible taxonomic collapsing and geological-timescale annotation of time-calibrated phylogenetic trees in R often require coordination among several packages and repeated code for label parsing, clade validation, plotting, and export. Workflow-managed analyses additionally benefit from non-interactive configuration, predictable diagnostics, and machine-readable exit status. Results: We present Rclade, an R package that consolidates the multi-package coordination required for taxonomic collapsing into a streamlined, single-function interface. Rclade provides (1) custom ggproto objects (GeomPolygonStraight/GeomSegmentStraight) that bypass coord_munch() interpolation to achieve straight-edge rendering of collapsed triangles in circular layouts; (2) automatic detection and parsing of four taxonomic-label formats (GTDB, Silva, NCBI, embedded) plus user-supplied custom regex, with explicit input-validation contracts and parsing-accuracy evaluation on real and derived test sets; and (3) workflow embeddability through YAML configuration, library-mode APIs, and standard Unix exit codes. Benchmarks on synthetic and real datasets (200-10,000 synthetic tips and real reference trees up to 10,122 tips; 5 replicates at every scale under a unified fully rendered measurement protocol) show that the full-pipeline overhead is modest for interactive use (median {approx}0.87 s in-session rendering and {approx}8.4 s process-level wall-clock at 10,000 tips). Conclusions: Rclade is a convenience layer over the ggtree/deeptime ecosystem that reduces boilerplate while adding targeted technical improvements for circular-layout rendering and format heterogeneity management.

5
Constructing microbiome co-occurrence networks with confidence: A conditional, nonparametric, inference-based approach

Song, H.; Xiang, Y.; Liu, H.; Ling, W.; Plantinga, A. M.; Srinivasan, S.; Dun, Y.; Zhao, N.; Sun, S.; Engel, S. M.; Simon, N.; Wu, M. C.

2026-09-01 bioinformatics 10.64898/2026.08.27.747483 medRxiv
Top 2%
0.8%
Show abstract

Constructing microbial association networks is a common strategy for exploring relationships among taxa in microbiome studies. Although marginal correlation methods are easy to implement and allow formal inference, they can produce spurious edges driven by indirect associations through other taxa. Conditional graphical-modeling methods aim to recover direct associations, but many rely on Gaussian or linear assumptions and often provide limited uncertainty quantification. We propose a conditional, nonparametric approach based on the scaled expected conditional covariance (SEcov). SEcov measures population-level conditional association by residualizing each taxon with respect to the remaining taxa and scaling the resulting expected conditional covariance. The resulting estimator can incorporate flexible machine-learning methods for conditional-mean estimation and admits asymptotic normal inference, enabling p-values and confidence intervals for taxon-pair associations. We demonstrate through simulation studies that our proposed approach improves network recovery relative to other methods, and we illustrate the new method via construction of a co-occurrence network for the vaginal microbiome during pregnancy. IMPORTANCEHigh-throughput sequencing has made it possible to characterize microbial communities at large scale, and network analysis is widely used to summarize relationships among taxa. However, networks based on marginal correlations may include indirect associations, whereas many conditional graphical models rely on assumptions that may be difficult to justify for sparse, zero-inflated, compositional microbiome data. SEcov offers a practical alternative by estimating conditional associations nonparametrically and attaching inferential uncertainty to individual edges. This allows investigators to construct microbiome networks using statistically interpretable evidence for taxon-pair associations, rather than relying solely on arbitrary correlation cutoffs or regularization tuning parameters.

6
Optimizing genomic selection: A comparison of SNP selection strategies for reduced-density panels in beef cattle

Ogunbawo, A. R.; Mulim, H. A.; Hidalgo, J.; Ventura, H. T.; Souza, N. O.; Oliveira, H. R.

2026-08-31 genetics 10.64898/2026.08.26.747408 medRxiv
Top 2%
0.5%
Show abstract

The exponential increase in the number of genotyped animals, combined with the availability of high-density SNP chips has introduced computational challenges for routine genomic evaluations, particularly during the construction of the genomic relationship matrix. Although higher-density SNP panels can facilitate the identification of causal mutations, their use substantially increases computational requirements without a proportional gain in genomic prediction performance. To optimize computational efficiency while maintaining accuracy of genomic predictions, this study compared five SNP selection strategies (i.e., random sampling, random sampling with inclusion of informative SNPs, linkage disequilibrium (LD)-based pruning, a Shannon entropy-based machine learning approach, and [[EQUATION]]-based prioritization) to develop reduced-density panels for Nellore cattle. Using high-density (HD) genotype data comprising 437,650 SNPs from 304,782 animals (after quality control) as reference, three reduced-density panels (25K, 45K, and 65K SNPs) panels were tested across five traits (i.e., Age at first calving, Stayability, Weaning weight, Yearling weight, Muscling) with diverse genetic architectures. Genomic estimated breeding values (GEBVs) derived from these reduced panels were compared to those obtained from the HD reference panel using Pearsons correlations, under both genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP) methods. In the GBLUP model, prediction accuracy generally improved with increased marker density. Random selection with and without the informative SNPs consistently yielded the highest accuracies, whereas the [[EQUATION]]-based approach showed the lowest agreement with the HD reference across all densities. In contrast, ssGBLUP demonstrated strong robustness to marker reduction, producing uniformly high correlations {approx}1.00) across all SNP densities and selection strategies. These findings indicate that optimized low-density SNP panels maintain prediction accuracy comparable to HD panels, offering a cost-effective tool for large-scale genomic evaluations.

7
Closing the biodiversity observation-to-action loop

Yamaguchi, K.; Uchida, K.; Hiraiwa, M.; Fukano, Y.

2026-08-31 ecology 10.64898/2026.08.27.747669 medRxiv
Top 3%
0.5%
Show abstract

Citizen science observations are abundant, but conservation requires turning uneven records into reliable predictions and directing new surveys to where information is missing. We developed a biodiversity platform for Japan that is updated monthly and integrates 2.32 million records to predict 8,297 species across seven taxonomic groups. Shared representation models outperformed species-specific models in four groups and extended predictions to species with few records. Five independent datasets, including structured monitoring, environmental DNA and complete forest inventories, confirmed that the models ranked observed species and occupied sites above alternatives, with median AUCs of 0.724 to 0.894 across sites and 0.650 to 0.841 across species. For any user-selected area, the platform returns candidate species, distribution predictions, a biodiversity map corrected for uneven observation effort, a conservation priority map for native species and a map recommending where to survey next. This map highlights places where species with few records are predicted to occur despite limited sampling. Independent observations showed that areas ranked highly by this predicted potential contained many such species, indicating that model predictions can help direct surveys toward knowledge gaps. New observations are incorporated into monthly updates, creating a national feedback system connecting citizen science, local conservation decisions and future surveys.

8
Patterns and Drivers of Diatom Diversity and Biogeography in the North Pacific

Barral, A.; Suzuki, K.; Kikuchi, Y.; Nakaoka, S.-i.; Takao, S.; Nakaoka, S.

2026-08-31 ecology 10.64898/2026.08.30.746603 medRxiv
Top 3%
0.4%
Show abstract

Marine diatoms contribute to about 20% of global primary production. We present the first basin-scale, multiyear assessment of diatom communities in the North Pacific, combining taxonomically high-resolution RuBisCO large subunit gene (rbcL) metabarcoding with concurrent environmental measurements. Using a nine-year time series of daily samples resolved at the species level via ~500 bp rbcL fragments, we performed multivariate analyses across biogeographic provinces, identifying significant correlations between community structure and environmental drivers such as temperature and macronutrient availability. We report the prevalence of a previously overlooked centric diatom species in the North Pacific, Eunotogramma lunatum, which appears to be near-dominant even in subarctic high-nitrate, low-chlorophyll waters where pennate diatoms are typically favored. These results demonstrate the power of rbcL for large-scale ocean monitoring and provide a critical baseline for future studies of diatom population dynamics, climate change impacts, and ecosystem resilience in a key marine region.

9
Automated wildlife re-identification by merging information from multiple body parts: A case study in sea turtles

Adam, L.; Montagna, M.; Roma, V.; Mancini, A.; Papafitsoros, K.

2026-08-31 ecology 10.64898/2026.08.28.747856 medRxiv
Top 4%
0.3%
Show abstract

Wildlife re-identification (re-ID) is a widely used and powerful tool with diverse applications in animal ecology and conservation. Current automated methods typically operate on single images of a single body part of the animal. However, a single encounter may contain multiple images capturing different body regions, each providing complementary individual-specific information. In contrast to automated approaches, researchers often manually select the most suitable images and regions for identification based on factors like visibility, occlusion and image quality. This creates a mismatch between automated methods and field practice, limiting the practical adoption of current automated re-ID pipelines. Here, we address this by introducing an encounter-based, multi-body-part re-ID framework, using sea turtles as a model taxon. Our framework combines three elements: (1) An orientation-aware deep learning model, TurtleDetector, that in addition to the full bodies, it also automatically segments key body regions, i.e. heads, front and hind flippers, from images within an encounter; (2) a hybrid body-part-specific retrieval method, that sequentially combines a fast global-feature model (MiewID or DINOv3) with a more accurate but costlier local-feature model (ALIKED with LightGlue); and (3) a merged identity-prediction strategy that selects the highest calibrated similarity score across all available body parts and images of an encounter. We evaluate the framework on three long-term re-ID datasets spanning three species, loggerheads, greens, and hawksbill turtles, under an evaluation protocol that mirrors real-world, time-aware re-ID workflows. Across datasets, combining multiple body regions consistently improved identification performance over the best-performing single body region, resulting to an increase of 4-6% in top-1 accuracy. Interestingly, body regions traditionally underused in sea turtle re-ID, such as the hind flippers and carapaces, provided complementary identifying information that improved encounter-level re-ID when integrated through the hybrid retrieval method. Our findings demonstrate that automated wildlife re-ID can benefit from moving beyond single-image, single-body-part identification towards encounter-level integration of all available visual evidence. Our work further suggests that, where feasible, field photo-acquisition protocols should aim to capture multiple informative views of an individual during each encounter. Importantly, many species and taxa, including elephants, primates, cetaceans, and other large vertebrates, possess such individual-specific features across multiple body regions, highlighting the broad potential applicability of our framework.

10
Insights for Estimating Animal Movement Step Selection Functions

Koshute, P.; Fagan, W. F.

2026-08-31 ecology 10.64898/2026.08.29.748012 medRxiv
Top 4%
0.3%
Show abstract

Ecologists remotely track movement steps of animals (e.g., via global positioning systems) and use step selection functions to study the effect of environmental factors upon their movement decisions. Constructing such functions requires pairing each observed step with some number of unobserved but feasible comparison steps. Larger numbers of comparison steps generally yield better estimates but also incur potentially challenging computational demands. Thus, it is important to determine an appropriate number of comparison steps. No established guidance exists for this decision. Here, we use simulated tracks to assess how many comparison steps are needed, fitting each set of steps to a conditional logistic regression model. We monitor errors in estimated effects for several classes of tracks, identifying the number of comparison steps for which mean relative absolute error in estimated effects is consistently low. By this criterion, 32 comparison steps per observed step are needed for our primary class of simulated tracks. Tracks in more homogeneous landscapes, tracks with shorter mean step lengths, or shorter tracks generally require more comparison steps (ranging from 64 to 128 per observed step) to achieve the same level of accuracy. Longer tracks generally require fewer comparison steps (16 per observed step). These results clearly demonstrate that the number of comparison steps influences how well step selection functions estimate covariate effects and provides initial direction in a research area that currently lacks quantitative guidance. Movement ecologists should take care when selecting the number of comparison steps paired with each observed step because those decisions matter.

11
TreeTOP: Plant experimental platforms in canopy space

Baumeister, J.; Bakhtiari, M. M.; Schreiber, M.; Eisenring, M.; Gossner, M.; Walden, S.; Becker, A.; Bouffaud, M. L.; Cesarz, S.; Dauphin, B.; Eisenhauer, N.; Goldmann, K.; Heidrich, L.; Jurburg, S.; Junker, R. R.; Kreuzwieser, J.; Lampei, C.; Nauss, T.; Peter, M.; Prada-Salcedo, L.; Tarkka, M.; Werner, C.; Zeuss, D.; Herrmann, S.; Buscot, F.; Heer, K.; Opgenoorth, L.

2026-08-31 ecology 10.64898/2026.08.30.748063 medRxiv
Top 4%
0.2%
Show abstract

1. Forest canopies harbour strong microclimatic gradients that shape plant performance, species interactions and ecosystem processes. Yet, despite renewed interest sparked by global change, forest canopies remain difficult-to-access experimental spaces. 2. With the goal to expand access to tree canopies as experimental arenas, we designed, built, and tested TreeTOP, a standardized experimental platform that opens canopy space for manipulative ecological experiments, specifically with potted plants. TreeTOP features lightweight aluminum frames placed in mature tree canopies non-invasively, allowing potted plants to be placed in three different heights, ground level, shade canopy, and sun canopy. 3. We implemented TreeTOP using two contrasting infrastructure concepts to demonstrate its applicability in both highly equipped canopy research facilities and forests without permanent canopy infrastructure. One installation relied on a canopy crane, grid power and fully automated irrigation, whereas the second was built by certified tree climbers and was equipped with an autonomous solar-powered, battery-operated irrigation system. At both sites, environmental sensor networks monitor the experiment. 4. TreeTOP successfully reproduced characteristic canopy microclimatic gradients, including increasing light availability, daytime air temperatures and thermal extremes with canopy height. Despite differing infrastructures, both implementations generated comparable microclimatic patterns, demonstrating that standardized canopy experiments are feasible in forests with or without permanent canopy access. By opening canopy space for manipulative experiments, TreeTOP provides a transferable framework for investigating plant performance, phenology, species interactions and microbiome assembly under realistic forest conditions.

12
Discovering 25 novel phyla that fill gaps in the eukaryotic tree of life

Tedersoo, L.; Mikryukov, V.; Sildever, S.; Chmolowska, D.; Piwosz, K.; Meyneng, M.; Monjot, A.; del Campo, J.; Lara, E.; Hakimzadeh, A.; Geisen, S.; Panksep, K.; Bahram, M.; Oliverio, A.; Shepherd, R.; Rückert, S.; Lanzen, A.; Hurdeal, V.; Concetta Eliso, M.; Casotti, R.; Hosseynimoghadam, M.; Siano, R.; Chauvet, M.; Prins, V.; Kisand, V.; Anslan, S.; Alkahtani, S.; Nilsson, H.

2026-08-31 microbiology 10.64898/2026.08.28.747736 medRxiv
Top 5%
0.2%
Show abstract

Protists play important roles in food chains and symbioses in soil and aquatic environments, displaying an enormous morphological and functional diversity. While most commonly found protist species are well known to science, our global-scale environmental DNA survey across soil, water, and sediments reveals dozens of novel, phylum-level phylogenetic lineages that remain to be characterized for basic morphology and function. A vast majority of these undescribed taxa occur in marine water and sediments, but some are common in soil. Most of these novel taxa have distinct substrate and habitat preferences and biogeographic patterns. To accord these lineages scientific agency and enable unambiguous scientific communication, we propose formal names for 150 species to phylum-level taxa from 25 deep lineages based on eDNA and rRNA gene long-read sequence information.

13
Design and Validation of New Primers for Specific and Sensitive Real-time PCR Detection and Quantification of Seven Botulinum Encoding Genes (Serotype A-G) of Clostridium botulinum

Phan, P.-L.; Chu, H.-A.; Le, T.-T.; Le, P.-A.; Nguyen, H.-L. T.; Tran, M.-N. T.; Nguyen, T.-T.; Pham, Y.; Phan, T.-N.

2026-09-01 molecular biology 10.64898/2026.08.21.746353 medRxiv
Top 5%
0.1%
Show abstract

Botulinum neurotoxins (BoNTs) comprise a highly diverse group of seven serotypes (from A-G) and over 40 subtypes worldwide. Previous primer- and probe-based nucleic acid amplification tests (NAATs) for detection of BoNT encoding genes are challenged by high levels of nucleotide polymorphism both across and within subtypes. In this study, multiple BoNT gene sequences were aligned to identify highly conserved regions for the design of new primers that enable the detection of all seven serotypes under the same conditions. Specific primer sets were designed and validated using in silico, conventional and real-time PCR with constructed plasmids carrying the target fragments and spiked food matrices. The established procedure achieved highly specific and sensitive detection of BoNT serotypes A-G with sensitivity of 10 copies/reaction and a total turnaround time of approximately 1.5 hours. The procedure also eliminated the carryover PCR product by using uracil-N-glycosylase in combination with dUTP in the assay reaction mix. This study provides an alternative NAAT with higher coverage and compliments the traditional mouse bioassays in enhancing global botulism surveillance capabilities.

14
Comprehensive study of Trypanosoma cruzi genetic diversity from Triatominae vectors in the Southern United States: Geographic structuring, mitochondrial introgression, and multiclonality

Hernandez, J. C.; Beatty, N. L.; Vogel, K. J.; Zima, J.; Novakova, E.

2026-08-31 microbiology 10.64898/2026.08.21.746190 medRxiv
Top 5%
0.1%
Show abstract

Background Trypanosoma cruzi, the causative agent of Chagas disease, is subdivided into distinct genetic groups known as Discrete Typing Units (DTUs), each with distinct genetic traits that influence epidemiology and transmission dynamics. Several triatomine species serve as potential vectors of T. cruzi in the United States. However, despite the growing number of Chagas disease cases in the country, little is known about the genetic diversity and population structure of T. cruzi in natural vector populations. Methodology/Principal Findings We applied a multilocus metabarcoding approach to improve DTU resolution and characterize the genetic diversity and structure of T. cruzi in triatomines collected across five states of the southern United States. Five single-copy nuclear markers and one mitochondrial marker were amplified and processed by high-throughput sequencing to assess genetic diversity. We recovered 35 nuclear and 15 mitochondrial haplotypes from 70 infected specimens. Overall, genetic diversity was low ({pi} < 0.01 at all nuclear loci), with DTUs TcI and the North American lineage of TcIV detected, TcI being the most prevalent. Geographic structuring was particularly evident in TcI strains, which exhibited a distinctive haplotype profile in Florida populations, potentially linked to the recently revalidated vector species Triatoma ambigua. Mitochondrial introgression from TcIV into TcI suggests inter-DTU genetic exchange in these populations. Multiple haplotypes within individual insects detected across single-copy nuclear markers, support multiclonal infection as common feature of T. cruzi in natural vectors. Conclusions/Significance These findings provide new insights into the genetic landscape and evolution of T. cruzi in the United States. Evolutionary connectivity through mitochondrial introgression and frequent multiclonality highlights the importance of deep sequencing approaches for resolving T. cruzi genetic diversity, with direct implications for understanding for transmission dynamics, disease monitoring and control.

15
PhageTAILor leverages machine learning for phage tail-like elements detection and classification in plant-associated bacteria

Cho, H.; Hour, S.; Roux, S.; Coclet, C.; Amusat, O.; Mutalik, V. K.; Kazakov, A. E.; Levy, A.; Nachmias, N.; Aureli, L.; Sweet, T. S.; Visel, A.; Ceballos, R. M.; Basso, J. T. R.

2026-09-01 microbiology 10.64898/2026.08.24.746745 medRxiv
Top 5%
0.1%
Show abstract

Phage tail-like elements (PTEs) -- tailocins, bacterial type VI secretion systems (T6SS), and extracellular contractile injection systems (eCIS) -- are contractile nanomachines that bacteria use to kill their neighbors and compete within their micro-ecosystems. PTEs help shape microbial community composition. Most PTE detection tools only detect a single PTE class. Moreover, most tailocin detection methods are largely restricted to Pseudomonas, leaving a key part of tailocin diversity uncharacterized. In this work, we present PhageTAILor (https://github.com/hjcho-bio/PhageTAILor), an integrative and fully automated pipeline that detects and classifies prophages and 3 PTE classes from bacterial genomes. PhageTAILor combines a 6-detector homology-based candidate search (geNomad, tail-gene, PHROGs-tail, SecReT6, eCIStem, and a divergence-tolerant tail-HMM detector) with a LightGBM classifier comprising 1 multiclass and 3 binary heads, trained on 6,501 bacterial genomes carrying 13,082 prophages and PTEs. A phylogeny-free feature matrix used in our model keeps predictions reproducible between model construction and user inference. PhageTAILor performs strongly at the genome level and generalizes beyond its Pseudomonas-rich training set. On a 76-strain cross-clade benchmark, PhageTAILor detected tailocins at F1 = 0.955. Furthermore, it identified 12 of 13 experimentally validated tailocins spanning five genera versus 2 of 13 for a Pseudomonas-restricted tool TattleTail. PhageTAILor also demonstrated sensitivity equivalent to viral detection tool geNomad while avoiding its higher false-positive rate. Applied to 7,925 plant- and soil-associated bacterial isolates, PhageTAILor showed that prophages in the phyllosphere and tailocins in plant-associated bacteria, whereas eCIS are enriched in soil. PhageTAILor is distributed as an open-source, modular pipeline with a command-line interface.

16
Pumping stations negatively affect the distribution of critically endangered European eel (Anguilla anguilla); a landscape-scale study using environmental DNA metabarcoding

Monaghan, A. I. T.; Griffiths, N. P.; Sellers, G. S.; Lawson Handley, L.; Nunn, A. D.; Hänfling, B.; Macarthur, J. A.; Wright, R. M.; Cattaneo, M.; Bolland, J. D.

2026-09-01 ecology 10.64898/2026.08.28.746846 medRxiv
Top 5%
0.1%
Show abstract

Context Pumping stations pose a threat to fish globally through land use change, habitat fragmentation and entrainment risk, with the catadromous and critically endangered European eel particularly impacted. Objectives/methods Establish, model, assess and understand the present-day distribution of European eel and resident fishes in 152 pumping station catchments in a once extensive wetland (The Fens) using eDNA metabarcoding (855 samples over two and half years), with specific focus on anthropogenic influences on hydrological connectivity and habitat quality. A removal survey design maximised confidence in negative results while minimising time and consumable costs. Results Eel occurrence upstream of pumping stations was low (occupancy = 28.3%) and positively associated with catchment area, fish species richness and natural hydrological connectivity (gravity drainage or flooding) and negatively associated with distance from the tidal limit. Fish species richness replaced catchment area and improved model performance, potentially acting as a biotic indicator of habitat quality and connectivity. Pumped catchments with manually operated upstream water transfers had reduced eel presence, potentially linked to the direction of water flow or the timing of operation. By contrast, fish species richness increased in these catchments during summer, suggesting displacement into unsuitable long-term habitats. Physical habitat maintenance had no detectable effect on eel occurrence or fish species richness. Conclusions This study provides the first landscape-scale assessment of European eel distribution and drivers of occurrence in pumped river catchments. The highly novel and comprehensive insights have implications for European eel conservation as well as infrastructure and catchment management, including compliance with legislation (EC Regulation No. 1100/2007).

17
Intelligent differential ion mobility spectrometry (iDMS): A deep neural network that predicts optimal space-resolved ion mobility parameters for isomeric monoglycosphingolipids

Nguyen-Tran, T.; Shi, X. X.; Hashimoto-Roth, E.; Organ, M. G.; Lavallee-Adam, M.; Perkins, T. J.; Bennett, S. A. L.

2026-09-01 bioinformatics 10.64898/2026.08.26.747394 medRxiv
Top 6%
0.1%
Show abstract

Simultaneous quantification of monoglycosphingolipid stereoisomers is required to monitor changes in defective enzymatic pathways linked to diseases such as Gaucher Disease, Parkinson's Disease, and Krabbe Disease. Resolution of beta-glucosyl and beta-galactosyl epimers cannot be achieved by standard liquid chromatography, electrospray ionization, tandem mass spectrometry (LC-ESI-MS/MS). Separation becomes possible when field asymmetric ion mobility spectrometry (FAIMS), also known as differential mobility mass spectrometry (DMS), is added as an orthogonal separation technique to LC. FAIMS/DMS separates epimeric ion clusters in a high versus low electric field (separation voltage, SV) then redirects the target epimeric ions to the mass spectrometer through the application of a direct current (compensation voltage, CoV). Resolving SVs and CoVs must be manually determined for each lipid. Manual derivation is a labour-intensive process that requires pure synthetic standards, limiting the number of stereoisomers a user can include in an assay. To address this problem, we introduce here intelligent DMS (iDMS). iDMS is an in silico supervised neural network model that learns the ion mobility relationships between SV and CoV and the monoglycosphingolipid structural features of sugar headgroup, N-acyl chain length, and N-acyl degree of unsaturation. iDMS predicts the SV and CoV combinations capable of resolving any stereoisomer pair from a training dataset of composed of measured signal intensities across a range of SVs and CoVs of 12 lipids. This machine learning alternative to manual DMS optimization promises to accelerate the deployment of multiple-reaction-monitoring mode (MRM) RPLC-ESI-DMS-MS/MS assays for the routine and rapid quantification of biologically relevant monoglycosphingolipid stereoisomers.

18
Secondary Structure Diversity of the Mitochondrial Small-Subunit rRNA in Porifera

Zhou, Y.; Gong, L.; Niu, G.; Shi, H.; Gutell, R.; Li, X.; Wei, M.

2026-08-30 evolutionary biology 10.64898/2026.08.28.747467 medRxiv
Top 6%
0.1%
Show abstract

Animal mitochondrial rRNAs are commonly viewed as structurally reduced, yet sponge mt SSU rRNAs range from compact to highly expanded structures. Using nine conserved structural anchors, we compared 216 taxonomically resolved records from four classes and 22 orders, including 16 freshwater Spongillida and 200 marine sponges. Twelve homologous hypervariable substructures were coded as structural types, and their ordered combinations as composite types. We identified 38 structural types and 62 composite types across molecules ranging from 853 to 2,019 nt. Hexactinellida and freshwater Spongillida were each uniform for a distinct composite type but differed markedly in overall structure: hexactinellid mt SSU rRNAs were compact, whereas those of Spongillida were long and contained four to five candidate insertion regions. These results show that a conserved scaffold can accommodate extensive lineage-associated structural variation and provide a practical framework for comparing highly divergent mitochondrial rRNAs.

19
eDNA reveals urban habitat-specific sorting of a mixed regional fish fauna into distinct biodiversity and life-history assemblages

Zapfe, K. L.; Parker, E.; Elias, D.; Hogue, G. M.; Dornburg, A.

2026-08-31 ecology 10.64898/2026.08.29.748007 medRxiv
Top 6%
0.1%
Show abstract

Urbanization is reshaping freshwater ecosystems, with well-documented effects across gradients of land-use change, hydrologic alteration, and habitat degradation. However, how biodiversity is organized among neighboring urban aquatic habitats that differ in hydrologic connectivity, disturbance transmission, residence time, management history, and opportunities for species movement is often less clear. This creates a challenge for interpreting urban fish communities at local scales as species occurrence may reflect both contemporary habitat filtering and historical contingencies including native persistence, interbasin transfer, stocking, and nonindigenous introductions. Here we use eDNA detections, historical records, phylogenetic information, and species trait data to investigate the fish assemblages of the Charlotte metropolitan region. We detect a highly mixed fauna that also depicts a strong signature of structured biodiversity profiles across taxonomic, phylogenetic, functional, and life-history dimensions between habitat types. In particular, bounded habitats contained assemblages with larger-bodied species that are fecund and faster to reproduce relative to free-flowing habitats. Species-level occurrence models did not support a simple trait-by-habitat rule. Instead our results demonstrate that urban aquatic habitats can sort historically mixed regional species pools into predictable assemblage-level life-history profiles while simultaneously retaining signatures of evolutionary and historical biogeographic contingency.

20
Accurate detection of metagenomic strain-level associations using average nucleotide identity with StrainSpy

Mallawaarachchi, S.; Tandon, K.; Rajan, N.; Marcelino, V. R.; Sandhu, S.; Bedoui, S.; Ingle, D. J.; Gunjur, A.; Tonkin-Hill, G.

2026-09-01 microbiology 10.64898/2026.08.30.748153 medRxiv
Top 6%
0.1%
Show abstract

Genetic variation among microbial strains of the same species can profoundly influence their phenotypes, ecological functions, and impacts on human health. Traditionally, the relative abundance of a species has been used to identify associations between the microbiome and disease. However, this approach overlooks intra-species genetic variation and is susceptible to spurious correlations arising from the compositional nature of abundance data and microbial load. Fast, k-mer-based algorithms can now accurately estimate strain-level Average Nucleotide Identity (ANI) in metagenomes. Despite its value as an orthogonal metric for strain-level analysis, methods for conducting ANI-based association studies remain limited. To address this, we developed StrainSpy, a statistical algorithm that identifies associations between containment ANI and variables of interest across a wide range of study designs, including longitudinal and multi-cohort designs. Re-analysis of a study examining gut microbiota recovery in 12 healthy adults following antibiotic exposure revealed novel strain-level associations, including a reduction in strain-level diversity despite species persistence. Applying StrainSpy to a multi-cohort analysis of 3,414 colorectal cancer metagenomes identified novel strain-level associations with colorectal cancer. However, in a separate collection of microbiome-immunotherapy studies, no individual strain was consistently associated across cohorts. Importantly, across both datasets, StrainSpy informed containment ANI-based machine learning models achieved comparable accuracy to traditional abundance-based methods. StrainSpy is publicly available as an R package github.com/gtonkinhill/strainspy.